Sequential Superparamagnetic Clustering for Unbiased Classification of High-Dimensional Chemical Data

نویسندگان

  • Thomas Ott
  • Albert Kern
  • Ansgar Schuffenhauer
  • Maxim Popov
  • Pierre Acklin
  • Edgar Jacoby
  • Ruedi Stoop
چکیده

For the clustering of chemical structures that are described by the Similog, ISIS count, and ISIS binary fingerprints, we propose a sequential superparamagnetic clustering approach. To appropriately handle nonbinary feature keys, we introduce an extension of the binary Tanimoto similarity measure. In our applications, data sets composed of structures from seven chemically distinct compound classes are evaluated and correctly clustered. The comparison, with results from leading methods, indicates the superiority of our sequential superparamagnetic clustering approach.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

High-Dimensional Unsupervised Active Learning Method

In this work, a hierarchical ensemble of projected clustering algorithm for high-dimensional data is proposed. The basic concept of the algorithm is based on the active learning method (ALM) which is a fuzzy learning scheme, inspired by some behavioral features of human brain functionality. High-dimensional unsupervised active learning method (HUALM) is a clustering algorithm which blurs the da...

متن کامل

Sequential clustering: tracking down the most natural clusters

Sequential superparamagnetic clustering (SSC) is a substantial extension of the superparamagnetic clustering approach (SC). We demonstrate that the novel method is able to master the important problem of inhomogeneous classes in the feature space. By fully exploiting the non-parametric properties of SC, the method is able to find the natural clusters even if they are highly different in shape a...

متن کامل

Clustered Multidimensional Scaling with Rulkov Neurons

When dealing with high-dimensional measurements that often show non-linear characteristics at multiple scales, a need for unbiased and robust classification and interpretation techniques has emerged. Here, we present a method for mapping high-dimensional data onto low-dimensional spaces, allowing for a fast visual interpretation of the data. Classical approaches of dimensionality reduction atte...

متن کامل

Supervised Feature Extraction of Face Images for Improvement of Recognition Accuracy

Dimensionality reduction methods transform or select a low dimensional feature space to efficiently represent the original high dimensional feature space of data. Feature reduction techniques are an important step in many pattern recognition problems in different fields especially in analyzing of high dimensional data. Hyperspectral images are acquired by remote sensors and human face images ar...

متن کامل

Optimal Feature Selection for Data Classification and Clustering: Techniques and Guidelines

In this paper, principles and existing feature selection methods for classifying and clustering data be introduced. To that end, categorizing frameworks for finding selected subsets, namely, search-based and non-search based procedures as well as evaluation criteria and data mining tasks are discussed. In the following, a platform is developed as an intermediate step toward developing an intell...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:
  • Journal of chemical information and computer sciences

دوره 44 4  شماره 

صفحات  -

تاریخ انتشار 2004